unified memory
Here's why Apple's Mac Studio has become so expensive
There's no question that Apple's Mac Mini is a great desktop computer, offering a surprising amount of power in a shockingly small size. However, it tops out at an M5 Pro processor and 64GB of RAM, with 307 GB/s of memory bandwidth. That's enough for many demanding situations, but AI and content creation power users may need more. The "base" version (starting at 2,500) has an M5 Max processor with an 18-core GPU and 32-core GPU, and up to 128GB of unified memory, depending on the option chosen. That goes up to 5,899 for the 18-core CPU and 40-core GPU version with 128GB of RAM and 2TB storage.
Apple Mac Studio M4 Max review: A creative powerhouse
The Mac Studio is Apple's ultimate performance computer, but this year's model came with a twist: It's equipped with either an M4 Max or an M3 Ultra processor. The latter might seem like a step backward, since nearly all Macs (except the Mac Pro) are now equipped with M4 chips. However, the M3 Ultra is indeed Apple's best-performing processor, which makes the new Mac Studio its fastest computer ever. While the M3 Ultra model appears highly capable for creative pros and engineers, it starts at 4,000 and goes way up from there. I'm intrigued by that model based on benchmarks I saw elsewhere, of course.
Inside Pascal: NVIDIA's Newest Computing Platform
Unlike other technical computing applications that require high-precision floating-point computation, deep neural network architectures have a natural resilience to errors due to the backpropagation algorithm used in their training. Storing FP16 data compared to higher precision FP32 or FP64 reduces memory usage of the neural network, allowing training and deployment of larger networks. Using FP16 computation improves performance up to 2x compared to FP32 arithmetic, and similarly FP16 data transfers take less time than FP32 or FP64 transfers. The GP100 SM ISA provides new arithmetic operations that can perform two FP16 operations at once on a single-precision CUDA Core, and 32-bit GP100 registers can store two FP16 values. Atomic memory operations are important in parallel programming, allowing concurrent threads to correctly perform read-modify-write operations on shared data.